<!DOCTYPE html>
<html class="client-nojs vector-feature-night-mode-disabled vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-1 vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-1 vector-sticky-header-enabled" lang="en" dir="ltr"><head>
<meta charset="UTF-8">
<title>Consensus CDS Project</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="canonical" href="https://en.wikipedia.org/wiki/Consensus_CDS_Project"> <link href="./mw/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/user.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link rel="stylesheet" type="text/css" href="./mw/site.styles.css">
<link rel="stylesheet" type="text/css" href="./mw/noscript.css">
<link rel="stylesheet" type="text/css" href="./footer.css">
<link rel="stylesheet" type="text/css" href="./vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Consensus_CDS_Project rootpage-Consensus_CDS_Project skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading">
<span id="openzim-page-title" class="mw-page-title-main"><span class="mw-page-title-main">Consensus CDS Project</span></span>
</h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="en" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="en" dir="ltr">
<p class="mw-empty-elt">
</p>
<style data-mw-deduplicate="TemplateStyles:r1295905060">
/* start https://en.wikipedia.org/ */
.mw-parser-output .infobox-subbox{padding:0;border:none;margin:-3px;width:auto;min-width:100%;font-size:100%;clear:none;float:none;background-color:transparent}.mw-parser-output .infobox-3cols-child{margin:auto}.mw-parser-output .infobox .navbar{font-size:100%}@media screen{html.skin-theme-clientpref-night .mw-parser-output .infobox-full-data:not(.notheme)>div:not(.notheme)[style]{background:#1f1f23!important;color:#f8f9fa}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .infobox-full-data:not(.notheme)>div:not(.notheme)[style]{background:#1f1f23!important;color:#f8f9fa}}@media(min-width:640px){body.skin--responsive .mw-parser-output .infobox-table{display:table!important}body.skin--responsive .mw-parser-output .infobox-table>caption{display:table-caption!important}body.skin--responsive .mw-parser-output .infobox-table>tbody{display:table-row-group}body.skin--responsive .mw-parser-output .infobox-table th,body.skin--responsive .mw-parser-output .infobox-table td{padding-left:inherit;padding-right:inherit}}
/* end https://en.wikipedia.org/ */
</style><table class="infobox vevent" style="width:"><caption class="infobox-title summary">CCDS Project</caption><tbody><tr><th colspan="2" class="infobox-header" style="background-color: lavender; background-color: light-dark(lavender,#353549) !important;">Content</th></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap">Description</th><td class="infobox-data">Convergence towards a standard set of gene annotations</td></tr><tr><th colspan="2" class="infobox-header" style="background-color: lavender; background-color: light-dark(lavender,#353549) !important;">Contact</th></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap"><a href="Research_center" class="mw-redirect" title="Research center">Research center</a></th><td class="infobox-data"><a href="National_Center_for_Biotechnology_Information" title="National Center for Biotechnology Information">National Center for Biotechnology Information</a><br><a href="European_Bioinformatics_Institute" title="European Bioinformatics Institute">European Bioinformatics Institute</a><br><a href="University_of_California%2C_Santa_Cruz" title="University of California, Santa Cruz">University of California, Santa Cruz</a><br><a href="Wellcome_Trust_Sanger_Institute" class="mw-redirect" title="Wellcome Trust Sanger Institute">Wellcome Trust Sanger Institute</a></td></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap">Authors</th><td class="infobox-data"><a href="Kim_D._Pruitt" title="Kim D. Pruitt">Kim D. Pruitt</a></td></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap">Primary citation</th><td class="infobox-data">Pruitt KD, et al (2009)<sup id="cite_ref-pmid19498102_1-0" class="reference"><a href="#cite_note-pmid19498102-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup></td></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap">Release date</th><td class="infobox-data">2009</td></tr><tr><th colspan="2" class="infobox-header" style="background-color: lavender; background-color: light-dark(lavender,#353549) !important;">Access</th></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap">Website</th><td class="infobox-data"><a rel="nofollow" class="external free" href="https://www.ncbi.nlm.nih.gov/projects/CCDS/CcdsBrowse.cgi">https://www.ncbi.nlm.nih.gov/projects/CCDS/CcdsBrowse.cgi</a></td></tr><tr><th colspan="2" class="infobox-header" style="background-color: lavender; background-color: light-dark(lavender,#353549) !important;">Miscellaneous</th></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap">Version</th><td class="infobox-data">CCDS Release 24</td></tr></tbody></table>
<p>The <b>Consensus Coding Sequence (CCDS) Project</b> is a collaborative effort to maintain a dataset of protein-coding regions that are identically <a href="DNA_annotation" title="DNA annotation">annotated</a> on the human and mouse reference genome assemblies. The CCDS project tracks identical protein annotations on the reference mouse and human genomes with a stable identifier (CCDS ID), and ensures that they are consistently represented by the National Center for Biotechnology Information <a href="National_Center_for_Biotechnology_Information" title="National Center for Biotechnology Information">(NCBI)</a>, <a href="Ensembl" class="mw-redirect" title="Ensembl">Ensembl</a>, and <a href="UCSC_Genome_Browser" title="UCSC Genome Browser">UCSC Genome Browser</a>.<sup id="cite_ref-pmid19498102_1-1" class="reference"><a href="#cite_note-pmid19498102-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup> The integrity of the CCDS dataset is maintained through stringent <a href="#Quality_assurance_testing">quality assurance testing</a> and on-going <a href="#Manual_curation">manual curation</a>.<sup id="cite_ref-Second_2-0" class="reference"><a href="#cite_note-Second-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup>
</p>
<meta property="mw:PageProp/toc">
<div class="mw-heading mw-heading2"><h2 id="Motivation_and_background">Motivation and background</h2></div>
<p>Biological and biomedical research has come to rely on accurate and consistent annotation of genes and their products on genome assemblies. Reference annotations of genomes are available from various sources, each with their own independent goals and policies, which results in some annotation variation.
</p><p>The CCDS project was established to identify a gold standard set of protein-coding gene annotations that are identically annotated on the human and mouse <a href="Reference_genome" title="Reference genome">reference genome</a> assemblies by the participating annotation groups. The CCDS gene sets that have been arrived at by consensus of the different partners <sup id="cite_ref-Second_2-1" class="reference"><a href="#cite_note-Second-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> now consist of over 18,000 human and over 20,000 mouse genes (see <a href="#CCDS_release_history">CCDS release history</a>). The CCDS dataset is increasingly representing more <a href="Alternative_splicing" title="Alternative splicing">alternative splicing</a> events with each new release.<sup id="cite_ref-third_3-0" class="reference"><a href="#cite_note-third-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Contributing_groups">Contributing groups</h2></div>
<p>Participating annotation groups include:<sup id="cite_ref-third_3-1" class="reference"><a href="#cite_note-third-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup>
</p>
<ul><li>National Center for Biotechnology Information <a href="National_Center_for_Biotechnology_Information" title="National Center for Biotechnology Information">(NCBI)</a></li>
<li>European Bioinformatics Institute <a href="European_Bioinformatics_Institute" title="European Bioinformatics Institute">(EBI)</a></li>
<li>Wellcome Trust Sanger Institute <a href="Wellcome_Trust_Sanger_Institute" class="mw-redirect" title="Wellcome Trust Sanger Institute">(WTSI)</a></li>
<li>HUGO Gene Nomenclature Committee <a href="HUGO_Gene_Nomenclature_Committee" title="HUGO Gene Nomenclature Committee">(HGNC)</a></li>
<li>Mouse Genome Informatics <a href="Mouse_Genome_Informatics" title="Mouse Genome Informatics">(MGI)</a></li></ul>
<p>Manual annotation is provided by:
</p>
<ul><li>Reference Sequence (<a href="RefSeq" title="RefSeq">RefSeq</a>) at NCBI</li>
<li>Human and Vertebrate Analysis and Annotation (HAVANA) at <a href="Wellcome_Trust_Sanger_Institute" class="mw-redirect" title="Wellcome Trust Sanger Institute">WTSI</a></li></ul>
<div class="mw-heading mw-heading2"><h2 id="Defining_the_CCDS_gene_set">Defining the CCDS gene set</h2></div>
<p>"Consensus" is defined as protein-coding regions that agree at the start codon, stop codon, and splice junctions, and for which the prediction meets quality assurance benchmarks.<sup id="cite_ref-pmid19498102_1-2" class="reference"><a href="#cite_note-pmid19498102-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup> A combination of manual and automated genome annotations provided by <a href="National_Center_for_Biotechnology_Information" title="National Center for Biotechnology Information">(NCBI)</a>
and <a href="Ensembl" class="mw-redirect" title="Ensembl">Ensembl</a> (which incorporates manual HAVANA annotations) are compared to identify annotations with matching genomic coordinates.
</p>
<div class="mw-heading mw-heading2"><h2 id="Quality_assurance_testing">Quality assurance testing</h2></div>
<p>In order to ensure that CDSs are of high quality, multiple quality assurance (QA) tests are performed (Table 1). All tests are performed following the annotation comparison step of each CCDS build and are independent of individual annotation group QA tests performed prior to the annotation comparison.<sup id="cite_ref-third_3-2" class="reference"><a href="#cite_note-third-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup>
</p>
<table class="wikitable">
<caption>Table 1: Examples of the types of CCDS QA tests performed prior to acceptance of CCDS candidates <sup id="cite_ref-third_3-3" class="reference"><a href="#cite_note-third-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup>
</caption>
<tbody><tr>
<th scope="col" width="290px">QA test
</th>
<th scope="col" width="640px">Purpose of the test
</th></tr>
<tr>
<td>Subject to NMD</td>
<td>Checks for transcripts that may be subject to nonsense-mediated decay (NMD)
</td></tr>
<tr>
<td>Low quality</td>
<td>Checks for low coding propensity
</td></tr>
<tr>
<td>Non-consensus splice sites</td>
<td>Checks for non-canonical splice sites
</td></tr>
<tr>
<td>Predicted pseudogene</td>
<td>Checks for genes that are predicted to be pseudogenes by UCSC
</td></tr>
<tr>
<td>Too short</td>
<td>Checks for transcripts or proteins that are unusually short, typically <100 amino acids
</td></tr>
<tr>
<td>Ortholog not found/not conserved</td>
<td>Checks for genes that are not conserved and/or are not in a HomoloGene cluster
</td></tr>
<tr>
<td>CDS start or stop not in alignment</td>
<td>Checks for a start or stop codon in the reference genome sequence
</td></tr>
<tr>
<td>Internal stop</td>
<td>Checks for the presence of an internal stop codon in the genomic sequence
</td></tr>
<tr>
<td>NCBI:Ensembl protein length different</td>
<td>Checks if the protein encoded by the NCBI RefSeq is the same length as the EBI/WTSI protein
</td></tr>
<tr>
<td>NCBI:Ensembl low percent identity</td>
<td>Checks for >99% overall identity between the NCBI and EBI/WTSI proteins
</td></tr>
<tr>
<td>Gene discontinued</td>
<td>Checks if the GeneID is no longer valid
</td></tr></tbody></table>
<p>Annotations that fail QA tests undergo a round of manual checking that may improve results or reach a decision to reject annotation matches based on QA failure.
</p>
<div class="mw-heading mw-heading2"><h2 id="Review_process">Review process</h2></div>
<p>The CCDS database is unique in that the review process must be carried out by multiple collaborators, and agreement must be reached before any changes can be made. This is made possible with a collaborator coordination system that includes a work process flow and forums for analysis and discussion. The CCDS database operates an internal website that serves multiple purposes including curator communication, collaborator voting, providing special reports and tracking the status of CCDS representations. When a collaborating CCDS group member identifies a CCDS ID that may need review, a voting process is employed to decide on the final outcome.
</p>
<div class="mw-heading mw-heading2"><h2 id="Manual_curation">Manual curation</h2></div>
<p>Coordinated manual curation is supported by a restricted-access website and a discussion e-mail list. CCDS curation guidelines were established to address specific conflicts that were observed at a higher frequency. Establishment of CCDS curation guidelines has helped to make the CCDS curation process more efficient by reducing the number of conflicting votes and time spent in discussion to reach a consensus agreement. A link to the CCDS curation guidelines can be found <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/CCDS/docs/CCDS_curation_guidelines.pdf">here</a>.
</p><p>Curation policies established for the CCDS data set have been integrated in to the <a href="RefSeq" title="RefSeq">RefSeq</a> and HAVANA annotation guidelines and thus, new annotations provided by both groups are more likely to be concordant and result in addition of a CCDS ID. These standards address specific problem areas, are not a comprehensive set of annotation guidelines, and do not restrict the annotation policies of any collaborating group.<sup id="cite_ref-Second_2-2" class="reference"><a href="#cite_note-Second-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> Examples include, standardized curation guidelines for selection of the initiation codon and interpretation of upstream <a href="Open_reading_frame" title="Open reading frame">ORFs</a> and transcripts that are predicted to be candidates for <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">nonsense-mediated decay</a>. Curation occurs continuously, and any of the collaborating centers can flag a CCDS ID as a potential update or withdrawal.
</p><p>Conflicting opinions are addressed by consulting with scientific experts or other annotation curation groups such as the HUGO Gene Nomenclature Committee <a href="HUGO_Gene_Nomenclature_Committee" title="HUGO Gene Nomenclature Committee">(HGNC)</a> and Mouse Genome Informatics <a href="Mouse_Genome_Informatics" title="Mouse Genome Informatics">(MGI)</a>. If a conflict cannot be resolved, then collaborators agree to withdraw the CCDS ID until more information becomes available.
</p>
<div class="mw-heading mw-heading2"><h2 id="Curation_challenges_and_annotation_guidelines">Curation challenges and annotation guidelines</h2></div>
<p><b>Nonsense-mediated decay (NMD):</b>
<a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a> is the most powerful <a href="Messenger_RNA" title="Messenger RNA">mRNA</a> surveillance process. <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a> eliminates defective <a href="Messenger_RNA" title="Messenger RNA">mRNA</a> before it can be translated into protein.<sup id="cite_ref-fourth_4-0" class="reference"><a href="#cite_note-fourth-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> This is important because if the defective <a href="Messenger_RNA" title="Messenger RNA">mRNA</a> is translated, the truncated protein may cause disease. Different mechanisms have been proposed to explain <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a>; one being the <a href="Exon_junction_complex" title="Exon junction complex">exon junction complex</a> (EJC) model. In this model, if the stop codon is >50 nt upstream of the last exon-exon junction, the transcript is assumed to be a <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a> candidate.<sup id="cite_ref-Second_2-3" class="reference"><a href="#cite_note-Second-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> The CCDS collaborators use a conservative method, based on the EJC model, to screen mRNA transcripts. Any transcripts determined to be <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a> candidates are excluded from the CCDS data set except in the following situations:<sup id="cite_ref-Second_2-4" class="reference"><a href="#cite_note-Second-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup>
</p>
<ol><li>all transcripts at one particular locus are assessed to be <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a> candidates however the locus is previously known to be protein coding region;</li>
<li>there is experimental evidence suggesting that a functional protein is produced from the <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a> candidate transcript.</li></ol>
<p>Previously, <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a> candidate transcripts were considered to be protein coding transcripts by both <a href="RefSeq" title="RefSeq">RefSeq</a> and HAVANA, and thereby, these <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a> candidate transcripts were represented in the CCDS data set. The <a href="RefSeq" title="RefSeq">RefSeq</a> group and the HAVANA project have subsequently revised their annotation policies.
</p><p><b>Multiple in-frame translation start sites:</b>
Multiple factors contribute to translation initiation, such as upstream <a href="Open_reading_frame" title="Open reading frame">open reading frames</a> (uORFs), secondary structure and the sequence context around the translation initiation site. A common start site is defined within Kozak consensus sequence: (GCC) GCCACCAUGG in vertebrates. The sequence in brackets (GCC) is the motif with unknown biological impact.<sup id="cite_ref-seventh_5-0" class="reference"><a href="#cite_note-seventh-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> There are variations within Kozak consensus sequence, such as G or A is observed three nucleotides upstream (at position -3) of AUG. Bases between positions -3 and +4 of Kozak sequence have the most significant impact on translational efficiency. Hence, a sequence (A/G)NNAUGG is defined as a strong Kozak signal in the CCDS project.
</p><p>According to the scanning mechanism, the small ribosomal subunit can initiate translation from the first reached start codon. There are exceptions to the scanning model:
</p>
<ol><li>when the initiation site is not surrounded by a strong Kozak signal, which results in leaky scanning. Thereby, the <a href="Ribosome" title="Ribosome">ribosome</a> skips this AUG and initiates translation from a downstream start site;</li>
<li>when a shorter <a href="Open_reading_frame" title="Open reading frame">ORF</a> can allow the <a href="Ribosome" title="Ribosome">ribosome</a> to re-initiate translation at a downstream <a href="Open_reading_frame" title="Open reading frame">ORF</a>.<sup id="cite_ref-seventh_5-1" class="reference"><a href="#cite_note-seventh-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup></li></ol>
<p>According to the CCDS annotation guidelines, the longest <a href="Open_reading_frame" title="Open reading frame">ORF</a> must be annotated except when there is experimental evidence that an internal start site is used to initiate translation. Additionally, other types of new data, such as ribosome profiling data,<sup id="cite_ref-Ninth_6-0" class="reference"><a href="#cite_note-Ninth-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup> can be used to identify start codons. The CCDS data set records one translation initiation site per CCDS ID. Any alternative start sites may be used for translation and will be stated in a CCDS public note.
</p><p><b>Upstream open reading frames:</b>
AUG initiation codons located within transcript leaders are known as upstream AUGs (uAUGs). Sometimes, uAUGs are associated with u<a href="Open_reading_frame" title="Open reading frame">ORFs</a> . u<a href="Open_reading_frame" title="Open reading frame">ORFs</a> are found in approximately 50% of human and mouse transcripts.<sup id="cite_ref-Sixth_7-0" class="reference"><a href="#cite_note-Sixth-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup> The existence of u<a href="Open_reading_frame" title="Open reading frame">ORFs</a> are another challenge for the CCDS data set. The scanning mechanism for translation initiation suggests that small ribosomal subunits (40S) bind at the 5’ end of a nascent <a href="Messenger_RNA" title="Messenger RNA">mRNA</a> transcript and scan for the first AUG start codon.<sup id="cite_ref-seventh_5-2" class="reference"><a href="#cite_note-seventh-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> It is possible that an uAUG is recognised first, and the corresponding uORF is then translated. The translated u<a href="Open_reading_frame" title="Open reading frame">ORF</a> could be a <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a> candidate, although studies have shown that some u<a href="Open_reading_frame" title="Open reading frame">ORFs</a> can avoid <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a>. The average size limit for u<a href="Open_reading_frame" title="Open reading frame">ORFs</a> that will escape <a href="Nonsense-mediated_decay" title="Nonsense-mediated decay">NMD</a> is approximately 35 <a href="Amino_acid" title="Amino acid">amino acids</a>.<sup id="cite_ref-Second_2-5" class="reference"><a href="#cite_note-Second-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Eighth_8-0" class="reference"><a href="#cite_note-Eighth-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup> It also has been suggested that u<a href="Open_reading_frame" title="Open reading frame">ORFs</a> inhibit translation of the downstream gene by trapping a <a href="Ribosome" title="Ribosome">ribosome</a> initiation complex and causing the <a href="Ribosome" title="Ribosome">ribosome</a> to dissociate from the <a href="Messenger_RNA" title="Messenger RNA">mRNA</a> transcript before it reaches the protein-coding regions.<sup id="cite_ref-fourth_4-1" class="reference"><a href="#cite_note-fourth-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Sixth_7-1" class="reference"><a href="#cite_note-Sixth-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup> Currently, no studies have reported the global impact of u<a href="Open_reading_frame" title="Open reading frame">ORFs</a> on translational regulation.
</p><p>The current CCDS annotation guidelines allow the inclusion of <a href="Messenger_RNA" title="Messenger RNA">mRNA</a> transcripts containing u<a href="Open_reading_frame" title="Open reading frame">ORFs</a> if they meet the following two biological requirements:<sup id="cite_ref-Second_2-6" class="reference"><a href="#cite_note-Second-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup>
</p>
<ol><li>the <a href="Messenger_RNA" title="Messenger RNA">mRNA</a> transcript has a strong Kozak signal;</li>
<li>the <a href="Messenger_RNA" title="Messenger RNA">mRNA</a> transcript is either ≥ 35 <a href="Amino_acid" title="Amino acid">amino acids</a> or overlaps with the primary <a href="Open_reading_frame" title="Open reading frame">open reading frame</a>.</li></ol>
<p><b>Read-through transcripts:</b>
Read-through transcripts are also known as <a href="Conjoined_gene" title="Conjoined gene">conjoined genes</a> or co-transcribed genes. Read-through transcripts are defined as transcripts combining at least part of one exon from each of two or more distinct known (partner) genes which lie on the same chromosome in the same orientation.<sup id="cite_ref-Tenth_9-0" class="reference"><a href="#cite_note-Tenth-9"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup> The biological function of read-through transcripts and their corresponding protein molecules remain unknown. However, the definition of a read-through gene in the CCDS data set is that the individual partner genes must be distinct, and the read-through transcripts must share ≥ 1 exon (or ≥ 2 splice sites except in the case of a shared terminal exon) with each of the distinct shorter loci.<sup id="cite_ref-Second_2-7" class="reference"><a href="#cite_note-Second-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> Transcripts are not considered to be read-through transcripts in the following circumstances:
</p>
<ol><li>when transcripts are produced from <a href="Overlapping_genes" class="mw-redirect" title="Overlapping genes">overlapping genes</a> but do not share same splice sites;</li>
<li>when transcripts are translated from genes that have nested structures relative to each other. In this instance, the CCDS collaborators and the <a href="HUGO_Gene_Nomenclature_Committee" title="HUGO Gene Nomenclature Committee">HGNC</a> have agreed that the read-through transcript be represented as a separate locus.</li></ol>
<p><b>Quality of reference genome sequence:</b>
As the CCDS data set is built to represent genomic annotations of human and mouse, the quality problems with the human and mouse <a href="Reference_genome" title="Reference genome">reference genome</a> sequences become another challenge. Quality problems occur when the reference genome is misassembled. Thereby the misassembled genome may contain premature <a href="Stop_codon" title="Stop codon">stop codons</a>, <a href="Frameshift_mutation" title="Frameshift mutation">frame-shift indels</a>, or likely polymorphic <a href="Pseudogene" title="Pseudogene">pseudogenes</a>. Once these quality problems are identified, the CCDS collaborators report the issues to the Genome Reference Consortium, which investigates and makes the necessary corrections.
</p>
<div class="mw-heading mw-heading2"><h2 id="Access_to_CCDS_data">Access to CCDS data</h2></div>
<p>The CCDS project is available from the NCBI CCDS data set page <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/projects/CCDS/CcdsBrowse.cgi">(here)</a>, which provides FTP download links and a query interface to acquire information about CCDS sequences and locations. CCDS reports can be obtained by using the query interface, which is located at the top of the CCDS data set page. Users can select various types of identifiers such as CCDS ID, gene ID, gene symbol, nucleotide ID and protein ID to search for specific CCDS information.<sup id="cite_ref-pmid19498102_1-3" class="reference"><a href="#cite_note-pmid19498102-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup> The CCDS reports (Figure 1) are presented in a table format, providing links to specific resources, such as a history report, <a href="Entrez" title="Entrez">Entrez Gene</a><sup id="cite_ref-Eleventh_10-0" class="reference"><a href="#cite_note-Eleventh-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup> or re-query the CCDS data set. The sequence identifiers table presents transcript information in <a href="Vertebrate_and_Genome_Annotation_Project" class="mw-redirect" title="Vertebrate and Genome Annotation Project">VEGA</a>, <a href="Ensembl" class="mw-redirect" title="Ensembl">Ensembl</a> and <a rel="nofollow" class="external text" href="https://web.archive.org/web/20100805211135/http://www.ncbi.nlm.nih.gov/sutils/blink.cgi?mode=query">Blink</a>. The chromosome location table includes the genomic coordinates for each individual exon of the specific coding sequence. This table also provides links to several different genome browsers, which allow you to visualise the structure of the coding region.<sup id="cite_ref-pmid19498102_1-4" class="reference"><a href="#cite_note-pmid19498102-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup> Exact nucleotide sequence and protein sequence of the specific coding sequence are also displayed in the section of CCDS sequence data.
</p>
<div class="mw-heading mw-heading2"><h2 id="Current_applications">Current applications</h2></div>
<p>The CCDS dataset is an integral part of the <a href="GENCODE" title="GENCODE">GENCODE</a> gene annotation project<sup id="cite_ref-Twelfth_11-0" class="reference"><a href="#cite_note-Twelfth-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup> and it is used as a standard for high-quality coding exon definition in various research fields, including clinical studies, large-scale <a href="Epigenomics" title="Epigenomics">epigenomic</a> studies, <a href="Exome" title="Exome">exome</a> projects and exon array design.<sup id="cite_ref-third_3-4" class="reference"><a href="#cite_note-third-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup> Due to the consensus annotation of CCDS exons by the independent annotation groups, <a href="Exome" title="Exome">exome</a> projects in particular have regarded CCDS coding exons as reliable targets for downstream studies (e.g., for <a href="Single-nucleotide_polymorphism" title="Single-nucleotide polymorphism">single nucleotide variant</a> detection), and these exons have been used as <a href="Coding_region" title="Coding region">coding region</a> targets in commercially available <a href="Exome" title="Exome">exome</a> kits.<sup id="cite_ref-Thirteenth_12-0" class="reference"><a href="#cite_note-Thirteenth-12"><span class="cite-bracket">[</span>12<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="CCDS_release_history">CCDS release history</h2></div>
<p>The CCDS data set size has continued to increase with both the computational genome annotation updates, which integrate new data sets submitted to the International Nucleotide Sequence Database Collaboration <a rel="nofollow" class="external text" href="http://www.insdc.org/">(INSDC</a>), and on ongoing curation activities that supplement or improve upon that annotation. Table 2 summarises the key statistics for each CCDS build where <b>Public CCDS IDs</b> are all those that were not under review or pending an update or withdrawal at the time of the current release date.
</p>
<table class="wikitable">
<caption>Table 2. Summary statistics for past CCDS releases.
</caption>
<tbody><tr>
<th scope="col" width="80px">Release
</th>
<th scope="col" width="150px">Species
</th>
<th scope="col" width="180px">Assembly name
</th>
<th scope="col" width="180px">Public CCDS ID count
</th>
<th scope="col" width="180px">Gene ID count
</th>
<th scope="col" width="180px">Current release date
</th></tr>
<tr>
<td>1</td>
<td><i>Homo sapiens</i></td>
<td>NCBI35</td>
<td>13,740</td>
<td>12,950</td>
<td>Mar 14, 2007
</td></tr>
<tr>
<td>2</td>
<td><i>Mus musculus</i></td>
<td>MGSCv36</td>
<td>13,218</td>
<td>13,012</td>
<td>Nov 28, 2007
</td></tr>
<tr>
<td>3</td>
<td><i>Homo sapiens</i></td>
<td>NCBI36</td>
<td>17,494</td>
<td>15,805</td>
<td>May 1, 2008
</td></tr>
<tr>
<td>4</td>
<td><i>Mus musculus</i></td>
<td>MGSCv37</td>
<td>17, 082</td>
<td>16,888</td>
<td>Jan 24, 2011
</td></tr>
<tr>
<td>5</td>
<td><i>Homo sapiens</i></td>
<td>NCBI36</td>
<td>19,393</td>
<td>17,053</td>
<td>Sep 2, 2009
</td></tr>
<tr>
<td>6</td>
<td><i>Homo sapiens</i></td>
<td>GRCh37</td>
<td>22,912</td>
<td>18,174</td>
<td>Apr 20, 2011
</td></tr>
<tr>
<td>7</td>
<td><i>Mus musculus</i></td>
<td>MGSCv37</td>
<td>21,874</td>
<td>19,507</td>
<td>Aug 14, 2012
</td></tr>
<tr>
<td>8</td>
<td><i>Homo sapiens</i></td>
<td>GRCh37.p2</td>
<td>25,354</td>
<td>18,407</td>
<td>Sep 6, 2011
</td></tr>
<tr>
<td>9</td>
<td><i>Homo sapiens</i></td>
<td>GRCh37.p5</td>
<td>26,254</td>
<td>18,474</td>
<td>Oct 25, 2012
</td></tr>
<tr>
<td>10</td>
<td><i>Mus musculus</i></td>
<td>GRCm38</td>
<td>22,934</td>
<td>19,945</td>
<td>Aug 5, 2013
</td></tr>
<tr>
<td>11</td>
<td><i>Homo sapiens</i></td>
<td>GRCh37.p9</td>
<td>27,377</td>
<td>18,535</td>
<td>Apr 29, 2013
</td></tr>
<tr>
<td>12</td>
<td><i>Homo sapiens</i></td>
<td>GRCh37.p10</td>
<td>27,655</td>
<td>18,607</td>
<td>Oct 24, 2013
</td></tr>
<tr>
<td>13</td>
<td><i>Mus musculus</i></td>
<td>GRCm38.p1</td>
<td>23,010</td>
<td>19,990</td>
<td>Apr 7, 2014
</td></tr>
<tr>
<td>14</td>
<td><i>Homo sapiens</i></td>
<td>GRCh37.p13</td>
<td>28,649</td>
<td>18,673</td>
<td>Nov 29, 2013
</td></tr>
<tr>
<td>15</td>
<td><i>Homo sapiens</i></td>
<td>GRCh37.p13</td>
<td>28,897</td>
<td>18,681</td>
<td>Aug 7, 2014
</td></tr>
<tr>
<td>16</td>
<td><i>Mus musculus</i></td>
<td>GRCm38.p2</td>
<td>23,835</td>
<td>20,079</td>
<td>Sep 10, 2014
</td></tr>
<tr>
<td>17</td>
<td><i>Homo sapiens</i></td>
<td>GRCh38</td>
<td>30,461</td>
<td>18,800</td>
<td>Sep 10, 2014
</td></tr>
<tr>
<td>18</td>
<td><i>Homo sapiens</i></td>
<td>GRCh38.p2</td>
<td>31,371</td>
<td>18,826</td>
<td>May 12, 2015
</td></tr>
<tr>
<td>19</td>
<td><i>Mus musculus</i></td>
<td>GRCm38.p3</td>
<td>24,834</td>
<td>20,215</td>
<td>July 30, 2015
</td></tr>
<tr>
<td>20</td>
<td><i>Homo sapiens</i></td>
<td>GRCh38.p7</td>
<td>32,524</td>
<td>18,892</td>
<td>Sep 8, 2016
</td></tr>
<tr>
<td>21</td>
<td><i>Mus musculus</i></td>
<td>GRCm38.p4</td>
<td>25,757</td>
<td>20,354</td>
<td>Dec 8, 2016
</td></tr>
<tr>
<td>22
</td>
<td><i>Homo sapiens</i>
</td>
<td>GRCh38.p12
</td>
<td>33,397
</td>
<td>19,033
</td>
<td>Jun 14, 2018
</td></tr>
<tr>
<td>23
</td>
<td><i>Mus musculus</i>
</td>
<td>GRCm38.p6
</td>
<td>27,219
</td>
<td>20,486
</td>
<td>Oct 24, 2019
</td></tr>
<tr>
<td>24
</td>
<td><i>Homo sapiens</i>
</td>
<td>GRCh38.p14
</td>
<td>35,608
</td>
<td>19,107
</td>
<td>Oct 26, 2022
</td></tr></tbody></table>
<p>The complete set of release statistics can be found at the official CCDS website on their <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/CCDS/CcdsBrowse.cgi?REQUEST=SHOW_STATISTICS#Current_Homo_sapiens_1">Releases & Statistics</a> page.
</p>
<div class="mw-heading mw-heading2"><h2 id="Future_prospects">Future prospects</h2></div>
<p>Long-term goals include the addition of attributes that indicate where transcript annotation is also identical (including the <a href="Untranslated_region" title="Untranslated region">UTRs</a>) and to indicate splice variants with different <a href="Untranslated_region" title="Untranslated region">UTRs</a> that have the same CCDS ID. It is also anticipated that as more complete and high-quality genome sequence data become available for other organisms, annotations from these organisms may be in scope for CCDS representation.
</p><p>The CCDS set will become more complete as the independent curation groups agree on cases where they initially differ, as additional experimental validation of weakly supported genes occurs, and as automatic annotation methods continue to improve. Communication among the CCDS collaborating groups is ongoing and will resolve differences and identify refinements between CCDS update cycles. Human updates are expected to occur roughly every 6 months and mouse releases yearly.<sup id="cite_ref-third_3-5" class="reference"><a href="#cite_note-third-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="See_also">See also</h2></div>
<ul><li><a href="GENCODE" title="GENCODE">GENCODE</a></li>
<li><a href="Human_Genome" class="mw-redirect" title="Human Genome">Human Genome</a></li>
<li><a href="Mouse_Genome_Informatics" title="Mouse Genome Informatics">Mouse Genome Informatics</a></li>
<li><a href="RefSeq" title="RefSeq">RefSeq</a></li>
<li><a href="Ensembl" class="mw-redirect" title="Ensembl">Ensembl</a></li></ul>
<div class="mw-heading mw-heading2"><h2 id="References">References</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239543626">
/* start https://en.wikipedia.org/ */
.mw-parser-output .reflist{margin-bottom:0.5em;list-style-type:decimal}@media screen{.mw-parser-output .reflist{font-size:90%}}.mw-parser-output .reflist .references{font-size:100%;margin-bottom:0;list-style-type:inherit}.mw-parser-output .reflist-columns-2{column-width:30em}.mw-parser-output .reflist-columns-3{column-width:25em}.mw-parser-output .reflist-columns{margin-top:0.3em}.mw-parser-output .reflist-columns ol{margin-top:0}.mw-parser-output .reflist-columns li{page-break-inside:avoid;break-inside:avoid-column}.mw-parser-output .reflist-upper-alpha{list-style-type:upper-alpha}.mw-parser-output .reflist-upper-roman{list-style-type:upper-roman}.mw-parser-output .reflist-lower-alpha{list-style-type:lower-alpha}.mw-parser-output .reflist-lower-greek{list-style-type:lower-greek}.mw-parser-output .reflist-lower-roman{list-style-type:lower-roman}
/* end https://en.wikipedia.org/ */
</style><div class="reflist">
<div class="mw-references-wrap mw-references-columns"><ol class="references">
<li id="cite_note-pmid19498102-1"><span class="mw-cite-backlink">^ <a href="#cite_ref-pmid19498102_1-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-pmid19498102_1-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-pmid19498102_1-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-pmid19498102_1-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-pmid19498102_1-4"><sup><i><b>e</b></i></sup></a></span> <span class="reference-text"><style data-mw-deduplicate="TemplateStyles:r1238218222">
/* start https://en.wikipedia.org/ */
.mw-parser-output cite.citation{font-style:inherit;word-wrap:break-word}.mw-parser-output .citation q{quotes:"\"""\"""'""'"}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}.mw-parser-output .id-lock-free.id-lock-free a{background:url("./mw/Lock-green.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-limited.id-lock-limited a,.mw-parser-output .id-lock-registration.id-lock-registration a{background:url("./mw/Lock-gray-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-subscription.id-lock-subscription a{background:url("./mw/Lock-red-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .cs1-ws-icon a{background:url("./mw/Wikisource-logo.svg")right 0.1em center/12px no-repeat}body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-free a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-limited a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-registration a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-subscription a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .cs1-ws-icon a{background-size:contain;padding:0 1em 0 0}.mw-parser-output .cs1-code{color:inherit;background:inherit;border:none;padding:inherit}.mw-parser-output .cs1-hidden-error{display:none;color:var(--color-error,#d33)}.mw-parser-output .cs1-visible-error{color:var(--color-error,#d33)}.mw-parser-output .cs1-maint{display:none;color:#085;margin-left:0.3em}.mw-parser-output .cs1-kern-left{padding-left:0.2em}.mw-parser-output .cs1-kern-right{padding-right:0.2em}.mw-parser-output .citation .mw-selflink{font-weight:inherit}@media screen{.mw-parser-output .cs1-format{font-size:95%}html.skin-theme-clientpref-night .mw-parser-output .cs1-maint{color:#18911f}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .cs1-maint{color:#18911f}}
/* end https://en.wikipedia.org/ */
</style><cite id="CITEREFPruittHarrowHarteWallin2009" class="citation journal cs1"><a href="Kim_D._Pruitt" title="Kim D. Pruitt">Pruitt KD</a>, Harrow J, Harte RA, Wallin C, Diekhans M, <a href="Donna_R._Maglott" title="Donna R. Maglott">Maglott DR</a>, Searle S, Farrell CM, Loveland JE, Ruef BJ, Hart E, Suner MM, Landrum MJ, Aken B, Ayling S, Baertsch R, Fernandez-Banet J, Cherry JL, Curwen V, Dicuccio M, Kellis M, Lee J, Lin MF, Schuster M, Shkeda A, Amid C, Brown G, Dukhanina O, Frankish A, Hart J, Maidak BL, Mudge J, Murphy MR, Murphy T, Rajan J, Rajput B, Riddick LD, Snow C, Steward C, Webb D, Weber JA, Wilming L, Wu W, Birney E, Haussler D, Hubbard T, Ostell J, Durbin R, Lipman D (2009). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2704439">"The consensus coding sequence (CCDS) project: Identifying a common protein-coding gene set for the human and mouse genomes"</a>. <i>Genome Res</i>. <b>19</b> (7): <span class="nowrap">1316–</span>23. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1101%2Fgr.080531.108">10.1101/gr.080531.108</a>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2704439">2704439</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/19498102">19498102</a>.</cite></span>
</li>
<li id="cite_note-Second-2"><span class="mw-cite-backlink">^ <a href="#cite_ref-Second_2-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Second_2-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-Second_2-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-Second_2-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-Second_2-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-Second_2-5"><sup><i><b>f</b></i></sup></a> <a href="#cite_ref-Second_2-6"><sup><i><b>g</b></i></sup></a> <a href="#cite_ref-Second_2-7"><sup><i><b>h</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFHarteFarrellLovelandSuner2012" class="citation journal cs1">Harte, RA; Farrell, CM; Loveland, JE; Suner, MM; Wilming, L; Aken, B; Barrell, D; Frankish, A; Wallin, C; Searle, S; Diekhans, M; Harrow, J; Pruitt, KD (2012). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3308164">"Tracking and coordinating an international curation effort for the CCDS project"</a>. <i>Database</i>. <b>2012</b>: bas008. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1093%2Fdatabase%2Fbas008">10.1093/database/bas008</a>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3308164">3308164</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/22434842">22434842</a>.</cite></span>
</li>
<li id="cite_note-third-3"><span class="mw-cite-backlink">^ <a href="#cite_ref-third_3-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-third_3-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-third_3-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-third_3-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-third_3-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-third_3-5"><sup><i><b>f</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFFarrellO'LearyHarteLoveland2014" class="citation journal cs1">Farrell, CM; O'Leary, NA; Harte, RA; Loveland, JE; Wilming, LG; Wallin, C; Diehans, M; Barrell, D; Searle, SM; Aken, B; Hiatt, SM; Frankish, A; Suner, MM; Rajput, B; Steward, CA; Brown, GR; Bennet, R; Murphy, M; Wu, W; Kay, MP; Hart, J; Rajan, J; Weber, J; Snow, C; Riddick, LD; Hunt, T; Webb, D; Thomas, M; Tamez, P; Rangwala, SH; McGarvey, KM; Pujar, S; Shkeda, A; Mudge, JM; Gonzale, JM; Gilbert, JG; Trevaion, SJ; Baetsch, R; Harrow, JL; Hubbard, T; Ostell, JM; Haussler, D; Pruitt, KD (2014). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3965069">"Current status and new features of the Consensus Coding Sequence database"</a>. <i>Nucleic Acids Res</i>. <b>42</b> (D1): <span class="nowrap">D865 –</span> <span class="nowrap">D872</span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1093%2Fnar%2Fgkt1059">10.1093/nar/gkt1059</a>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3965069">3965069</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/24217909">24217909</a>.</cite></span>
</li>
<li id="cite_note-fourth-4"><span class="mw-cite-backlink">^ <a href="#cite_ref-fourth_4-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-fourth_4-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFAlbertsJohnsonLewisRaff2002" class="citation book cs1">Alberts, B; Johnson, A; Lewis, J; Raff, M; Roberts, K; Walter, P (2002). <i>Molecular Biology of the Cell 5th edn</i>. New York: Garland Science.</cite></span>
</li>
<li id="cite_note-seventh-5"><span class="mw-cite-backlink">^ <a href="#cite_ref-seventh_5-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-seventh_5-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-seventh_5-2"><sup><i><b>c</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFKozak2002" class="citation journal cs1">Kozak, M (2002). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7126118">"Pushing the limits of the scanning mechanism for initiation of translation"</a>. <i>Gene</i>. <b>299</b> (<span class="nowrap">1–</span>2): <span class="nowrap">1–</span>34. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2FS0378-1119%2802%2901056-9">10.1016/S0378-1119(02)01056-9</a>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7126118">7126118</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/12459250">12459250</a>.</cite></span>
</li>
<li id="cite_note-Ninth-6"><span class="mw-cite-backlink"><b><a href="#cite_ref-Ninth_6-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFIngoliaBrarRouskinMcGeachy2014" class="citation journal cs1 cs1-prop-long-vol">Ingolia, NT; Brar, GA; Rouskin, S; McGeachy, AM; Weissman, JS (2014). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3775365">"Genome-wide Annotation and Quantitation of Translation by Ribosome Profiling"</a>. <i>Curr. Protoc. Mol. Biol</i>. Chapter 4: 4.18.1–4.18.19. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1002%2F0471142727.mb0418s103">10.1002/0471142727.mb0418s103</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>9780471142720</bdi>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3775365">3775365</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/23821443">23821443</a>.</cite></span>
</li>
<li id="cite_note-Sixth-7"><span class="mw-cite-backlink">^ <a href="#cite_ref-Sixth_7-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Sixth_7-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFCalvoPagliarniMootha2009" class="citation journal cs1">Calvo, SE; Pagliarni, DJ; Mootha, VK (2009). <a rel="nofollow" class="external text" href="http://dspace.mit.edu/bitstream/1721.1/50259/1/Calvo-2009-Upstream%20open%20readin.pdf">"Upstream open reading frames cause widespread reduction of protein expression and are polymorphic among humans"</a> <span class="cs1-format">(PDF)</span>. <i>Proc. Natl. Acad. Sci. U.S.A</i>. <b>106</b> (18): <span class="nowrap">7507–</span>12. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2009PNAS..106.7507C">2009PNAS..106.7507C</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1073%2Fpnas.0810916106">10.1073/pnas.0810916106</a></span>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2669787">2669787</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/19372376">19372376</a>.</cite></span>
</li>
<li id="cite_note-Eighth-8"><span class="mw-cite-backlink"><b><a href="#cite_ref-Eighth_8-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFSilvaPereiraMorgadoKong2006" class="citation journal cs1">Silva, AL; Pereira, FJC; Morgado, A; Kong, J; Martins, R; Faustino, P; Liebhaber, SA; Romao, L (2006). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC1664719">"The canonical UPF1-dependent nonsense-mediated mRNA decay is inhibited in transcripts carrying a short open reading frame independent of sequence context"</a>. <i>RNA</i>. <b>12</b> (12): <span class="nowrap">2160–</span>70. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1261%2Frna.201406">10.1261/rna.201406</a>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC1664719">1664719</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/17077274">17077274</a>.</cite></span>
</li>
<li id="cite_note-Tenth-9"><span class="mw-cite-backlink"><b><a href="#cite_ref-Tenth_9-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFPrakashSharmaAdatiOzawa2010" class="citation journal cs1">Prakash, Tulika; Sharma, Vineet K.; Adati, Naoki; Ozawa, Ritsuko; Kumar, Naveen; Nishida, Yuichiro; Fujikake, Takayoshi; Takeda, Tadayuki; Taylor, Todd D.; Michalak, Pawel (12 October 2010). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2953495">"Expression of Conjoined Genes: Another Mechanism for Gene Regulation in Eukaryotes"</a>. <i>PLOS ONE</i>. <b>5</b> (10): e13284. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2010PLoSO...513284P">2010PLoSO...513284P</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1371%2Fjournal.pone.0013284">10.1371/journal.pone.0013284</a></span>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2953495">2953495</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/20967262">20967262</a>.</cite></span>
</li>
<li id="cite_note-Eleventh-10"><span class="mw-cite-backlink"><b><a href="#cite_ref-Eleventh_10-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFMaglottOstellPruittTatusova2010" class="citation journal cs1"><a href="Donna_R._Maglott" title="Donna R. Maglott">Maglott, D.</a>; Ostell, J.; Pruitt, K. D.; Tatusova, T. (28 November 2010). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3013746">"Entrez Gene: gene-centered information at NCBI"</a>. <i>Nucleic Acids Res</i>. <b>39</b> (Database): <span class="nowrap">D52 –</span> <span class="nowrap">D57</span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1093%2Fnar%2Fgkq1237">10.1093/nar/gkq1237</a>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3013746">3013746</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/21115458">21115458</a>.</cite></span>
</li>
<li id="cite_note-Twelfth-11"><span class="mw-cite-backlink"><b><a href="#cite_ref-Twelfth_11-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFHarrowFrankishGonzalezTapanari2012" class="citation journal cs1">Harrow, J.; Frankish, A.; Gonzalez, J. M.; Tapanari, E.; Diekhans, M.; Kokocinski, F.; Aken, B. L.; Barrell, D.; Zadissa, A.; Searle, S.; Barnes, I.; Bignell, A.; Boychenko, V.; Hunt, T.; Kay, M.; Mukherjee, G.; Rajan, J.; Despacio-Reyes, G.; Saunders, G.; Steward, C.; Harte, R.; Lin, M.; Howald, C.; Tanzer, A.; Derrien, T.; Chrast, J.; Walters, N.; Balasubramanian, S.; Pei, B.; Tress, M.; Rodriguez, J. M.; Ezkurdia, I.; van Baren, J.; Brent, M.; Haussler, D.; Kellis, M.; Valencia, A.; Reymond, A.; Gerstein, M.; Guigo, R.; Hubbard, T. J. (5 September 2012). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3431492">"GENCODE: The reference human genome annotation for The ENCODE Project"</a>. <i>Genome Res</i>. <b>22</b> (9): <span class="nowrap">1760–</span>1774. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1101%2Fgr.135350.111">10.1101/gr.135350.111</a>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3431492">3431492</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/22955987">22955987</a>.</cite></span>
</li>
<li id="cite_note-Thirteenth-12"><span class="mw-cite-backlink"><b><a href="#cite_ref-Thirteenth_12-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFParlaIossifovGrabillSpector2011" class="citation journal cs1">Parla, Jennifer S; Iossifov, Ivan; Grabill, Ian; Spector, Mona S; Kramer, Melissa; McCombie, W Richard (2011). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3308060">"A comparative analysis of exome capture"</a>. <i>Genome Biol</i>. <b>12</b> (9): R97. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1186%2Fgb-2011-12-9-r97">10.1186/gb-2011-12-9-r97</a></span>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3308060">3308060</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/21958622">21958622</a>.</cite></span>
</li>
</ol></div></div>
<div class="mw-heading mw-heading2"><h2 id="External_links">External links</h2></div>
<ul><li><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/projects/CCDS/CcdsBrowse.cgi">CCDS home page</a></li></ul></div><!--htdig_noindex--><div><div class="zim-footer">
This article is issued from <a class="external text" title="Last edited on 2025-07-19" href="https://en.wikipedia.org/wiki/?title=Consensus_CDS_Project&oldid=1301317428">Wikipedia</a>. The text is available under <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">Creative Commons Attribution-Share Alike 4.0</a> unless otherwise noted. Additional terms may apply for the media files.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>
</body></html>